NSF PAR Search | NSF Public Access Repository

https://doi.org/10.1145/3764944.3764952

Han, Bing-Shiun; Parekh, Kunaal; Lin, Wan-Chu; Paul, Tathagata; Gandhi, Anshul; Liu, Zhenhua (August 2025, ACM SIGMETRICS Performance Evaluation Review)

GPU sharing between workloads is an e!ective approach to increase GPU utilization and reduce idle power waste. To minimize resource contention under GPU sharing, current architectures allow users to allocate core GPU compute resources exclusively to workloads. However, identifying the most e''cient GPU compute resource allocation for colocated workloads is challenging, as it requires balancing potential performance degradation and power savings. This paper presents a framework for finding the most energy-e''cient compute allocation for colocated workload pairs under NVIDIA MPS using lightweight prediction models. Experimental results, using a range of training, inference, and general CUDA workloads, demonstrate that our solution outperforms the equal sharing strategy by 35%, on average, and is within 1.5% of the o#ine optimal strategy.

Full Text Available

Search for: All records